The LLM judge is itself a measurement instrument (§11.2), so its own reliability is a defined quality gate — measured against an expert-labelled gold set.
0.74
Cohen's κ
judge ↔ expert
Above the κ ≥ 0.70 required before the judge gates anything. Below that → rubric re-calibration, not deployment.
Per-dimension agreement · accuracy vs gold set
Causal validity is the weakest — the same dimension below its programme target. Re-calibration focus.
Re-validation
last validated2026-06-09
rubricicam-v1.3 locked
judge modelclaude · cortex AI_COMPLETE
gold set128 expert-labelled
triggeron rubric / model change
Re-validated whenever the rubric_catalog content-hash or judge model changes. A standing human-audit sample detects drift and entrenched blind spots.
Societal / ethical risk (§11.4)
Operator-blame ratio
organisational-factor vs individual-action findings
1.6 : 1
Monitored so the rubric doesn't entrench systemic operator-blame bias. A ratio skewing toward individual actions would flag the judge, not the workforce.
Assist, don't decide
Every flagged case routes to a qualified HSE reviewer who decides assure / return-for-rework. The harness triages and explains with ICAM-layer + evidence rationale — it never closes an investigation. Safety-critical and regulated: human-in-the-loop is mandatory.